Forensic Science International: Genetics
○ Elsevier BV
Preprints posted in the last 90 days, ranked by how well they match Forensic Science International: Genetics's content profile, based on 26 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Alsuwaidi, M. S.; Albastaki, A.; Almulla, H.; Omar, A. K.; Almarri, M. A.
Show abstract
Short tandem repeat (STR) profiling is the cornerstone of forensic DNA analysis, traditionally performed via capillary electrophoresis. Recently, next generation sequencing has gained prominence due to its increased discriminatory power and enhanced performance with degraded samples. Nanopore sequencing offers a portable and cost-effective alternative, however historically high error rates have precluded its forensic adoption. Here, we evaluate R10.4.1 flow cell chemistry and multiple basecalling tiers (HAC, SUP, HYP) across several iterations (v4.2, v5.0, v5.2, v6.0) to assess their impact on genotyping accuracy. Analyzing 45 STR loci (22 autosomal and 23 Y-STRs) across single-source controls, we introduce a parallelized, user-friendly pipeline designed to transform raw POD5 files into STR profiles. Our results demonstrate a progressive improvement in genotyping accuracy with each basecalling iteration, with the latest models achieving 99.0% autosomal and 100% Y-STR concordance. Furthermore, we find that filtering on raw-read quality scores significantly improves genotyping by reducing background noise and generating cleaner profiles. Notably, the HYPv5.0 Q20 filter drove an average 53.2% reduction in misaligned reads across all loci in comparison to earlier basecalling models. Our study demonstrates that continual bioinformatic improvements in basecalling models, coupled with R10.4.1 chemistry, can provide accurate STR profiles in single-source samples, warranting larger validation studies with more diverse samples to further evaluate performance.
Berger, J.; Krawczak, M.; Zandstra, D.; Kayser, M.; Ralf, A.; Scheurer, E.; Caliebe, A.; Schulz, I.
Show abstract
The formal assessment of a genetic match between a suspect and some biological trace material is one of the key tasks of forensic genetics, particularly in cases of sexual offence. The analysis of Y-chromosomal short tandem repeats (Y-STRs) has proven especially useful in this context. For a long time, however, calculating the probability of a perfect Y-STR profile match under the defense hypothesis that the suspect was not the trace donor posed a great challenge. This was due to the inherent uncertainty about the population of alternative donors, the so-called suspect population. We recently proposed to resolve this controversy by systematically favoring the suspect and considering his close male relatives as the suspect population. However, since the mathematical framework developed for this purpose was simulation-based, its practical application turned out increasingly difficult with increasing pedigree size. Here, we present an adaptation of the so-called Elston-Stewart algorithm, originally developed for the linkage analysis of human genetic diseases, to allow calculation of exact match probabilities in a time that scales linearly with pedigree size. The adapted algorithm was implemented in a publicly available software tool, and its correctness was verified by the comparison of its output with the correct, analytical results obtained for selected example pedigrees. The new implementation mostly outperforms the simulation-based solution, albeit with the important exception of Y-STRs present in multiple copies. Given the increasingly prominent role of such multicopy markers in forensic genetics, the complementary use of both approaches appears the most sensible strategy for the time being.
Ismach, H.; Greenbaum, G.; Kennett, D.; Carmi, S.
Show abstract
Forensic investigative genetic genealogy (FIGG) is a revolutionary method in forensic genetics, whereby genetic relatives of an unknown target person are detected in direct-to-consumer genomic databases; the family trees of these relatives are then reconstructed to suggest candidates for the target person. Despite recent successes, the full potential of the technology has not yet been systematically evaluated at the scale of an entire country. To estimate the proportion of FIGG cases where a genetic relative is detected in the database (the "match rate"), we used the Israeli population register, covering all past and present citizens. After extensive quality control, the register included 12.4 million individuals, among them 10.4 million alive. We simulated genomic databases by randomly selecting subsets of predefined sizes of the live adult population. With a database covering 1% of the population, we estimate that FIGG would detect at least one relative of fifth degree (e.g., a second cousin) or closer for about 25% of the population, and at least two relatives for 9% of the population. A database covering 5% of the population would find at least one relative of third degree (e.g., a first cousin) or closer for half the population. More distant relatives are rarely identified in the register, likely due to its limited time depth. Our results provide the first country-wide direct estimates for the utility of FIGG in generating investigative leads.
Poggiali, B.; Aagreen, C. I. V.; Meyer, O. L.; Jepsen, A. H.; Korneliussen, T. S.; Kampmann, M.-L.; Borsting, C.; Andersen, J. D.
Show abstract
Shotgun sequencing (SGS) enables simultaneous interrogation of a broad range of loci across the human genome, even from low-template and highly degraded DNA samples. While human identification traditionally relies on short tandem repeats (STRs) due to their high polymorphism, standard forensic STRs (100-450 bp) are poorly suited for the short read ([~]150 bp) constraint of SGS. The purpose of this study was to evaluate the analysis limitations of standard forensic STRs in SGS data and to identify a novel panel of STRs optimised for short-read genomic data. First, we benchmarked four STR genotyping software tools (STRait Razor, GangSTR, STRinNGS, and HipSTR) by analysing 53 standard forensic STRs in SGS data. HipSTR showed the best performance but achieved only a call rate of 64.5% and an accuracy of 83.8%, and its performance was strongly affected by STR allele length and read depth. To overcome these constraints, we screened the population-wide 1000 Genomes Project dataset and identified a panel of 265 autosomal ultra-short (< 50 bp) STRs with an effective number of alleles (Ae) ranging from 3.0 to 7.5. As few as seven of these loci were sufficient to achieve a Mean Match Probability (MMP) below 1 x 10-6. To validate these findings, we developed a custom PCR-based amplicon sequencing panel targeting 97 of the most polymorphic ultra-short STRs and evaluated these in 41 blood samples from Danish individuals. The polymorphic nature of the selected loci was confirmed (Aeranged from 2.4 to 7.2). Our results furthermore demonstrated high concordance between the amplicon panel and SGS-derived genotypes, which substantiates that these ultra-short STRs provide a robust and highly polymorphic alternative for human identification in SGS data. Author summaryShotgun sequencing (SGS) methods are increasingly being adopted in fields such as forensic genetics. SGS yields large amounts of genetic information by reading short fragments across the entire genome, enabling a wide range of analyses that may be exploited as leads in a police investigation. Human identification has traditionally been based on STR loci with a PCR amplicon length of 100-450 base pairs. However, these loci are often longer than the reads generated by SGS data, which makes them difficult to analyse in a reliable way. In this study, we evaluated four software tools designed to genotype STRs and confirmed the limited ability to genotype traditional forensic STRs in SGS data. To address this limitation, we identified a new set of highly polymorphic ultra-short STRs (less than 50 base pairs in length) that enable robust human identification using SGS data. Despite their shorter length, these loci retain the multi-allelic nature inherent to traditional STRs. This ensures a low random match probability that is comparable with the standard forensic STR panels. The ultra-short STRs may be genotyped from highly degraded DNA and may provide the possibility for complex mixture analysis and multi-donor deconvolution, which makes the STRs uniquely suited for forensic casework.
Sümer, A. P.; Iasi, L. N. M.; Bossoms Mesa, A.; Slon, V.; Essel, E.; Hajdinjak, M.; Zorn, J.; Schmidt, A.; Nagel, S.; Nickel, B.; Viola, B.; Ziganshin, R.; Buzhilova, A.; Derevianko, A.; Pääbo, S.; Peter, B. M.
Show abstract
The Teshik-Tash 1 child whose remains were found in Uzbekistan represents the southeastern-most extent of the known Neandertal range, providing an important link with the better studied Caucasus and Altai Mountain ranges. However, due to poor DNA preservation, studying the genetics of Teshik Tash 1 has remained elusive. Here we present analyses of the nuclear DNA from the Teshik-Tash 1, from extracts that are highly contaminated with present-day human DNA. To achieve this, we developed a new computational method, admixslug, that jointly models contamination and population relationships, in order to infer the relationship of a target individual from which only low-quality nuclear DNA is available, to high-quality archaic human genomes. After validating admixslug, we show that Teshik-Tash 1 is genetically more similar to later Neandertals from Western Eurasia than to older Neandertals from the Altai Mountains. We estimate that Teshik-Tash 1 split from the Western Eurasian lineage between 80,000 and 100,000 years ago. Despite the geographical proximity of Teshik-Tash 1 to the Denisovan range, we find no evidence for Denisovan ancestry in his genome. Our results demonstrate that admixslug enables the study of archaic human specimens in cases where DNA preservation was previously considered too poor for population genetic analyses.
Walinjkar, A.
Show abstract
Background: Circulating tumour DNA (ctDNA) liquid biopsy is now established across oncology for early cancer detection, minimal residual disease surveillance, and treatment monitoring. Detection thresholds for all current ctDNA assays are derived empirically through receiver operating characteristic analysis on training cohorts - a statistically valid but theoretically uninformed approach that does not specify the minimum detectable tumour fraction given assay technical characteristics, nor identify when increasing sequencing depth ceases to provide additional clinical information. Methods: We model ctDNA detection as a binary hypothesis testing problem with Binomial-distributed mutant allele counts against a sequencing error noise floor. The Neyman-Pearson lemma is applied to derive the uniformly most powerful detector and the minimum detectable tumour fraction in closed form. The sequencing assay is modelled as a binary symmetric channel and Shannon channel capacity is calculated. Empirical validation uses n=61 data points extracted from five published peer-reviewed analytical validation studies across five independent institutions in the US and EU (2018 - 2025): Yu et al. 2022, Stetson et al. 2018, Frydendahl et al. 2023, Northcott et al. 2024, and Cheng et al. 2025. Results: The minimum detectable tumour fraction is derived in closed form as f_min approximately equal to (z_alpha + z_beta) multiplied by the square root of (epsilon divided by N), where N is sequencing depth, epsilon is the platform error rate, and z_alpha, z_beta are standard normal quantiles at the specified false positive and false negative rates. Shannon channel capacity is C = 1 minus H(epsilon) bits per read, where H(epsilon) is binary entropy. Empirical validation yields 84.3% agreement for single-locus assays. Discordance for multi-locus tumour-informed assays (NeXT Personal, duplex WGS) is consistent with the single-locus model scope and identifies the principal theoretical extension required. Conclusions: This framework provides the first formal Neyman-Pearson optimality proof for ctDNA detection, a closed-form detection limit, and a platform-independent efficiency metric for NHS and regulatory standardisation. Keywords: circulating tumour DNA; liquid biopsy; Neyman-Pearson detection; Shannon channel capacity; sequencing depth; limit of detection; minimal residual disease; signal detection theory
De Barba, M.; Boyer, F.; Baur, M.; Konec, M.; Pazhenkova, E.; Remollino, N.; Stoffel, C.; Boljte, B.; Miquel, C.; Skrbinsek, T.; Taberlet, P.; Fumagalli, L.
Show abstract
High-throughput amplicon sequencing has transformed microsatellite (STR) genotyping by overcoming many of the limitations of fragment-length analysis, enabling more accurate, cost-effective, and standardized genotyping. Yet, protocols specifically designed for high-throughput sequencing (HTS)-based STR genotyping from low-template and degraded DNA remain scarce, despite the prevalence of these challenging sample types in ecological and conservation contexts. We present a methodology for the de novo development of robust STR multiplex panels together with a laboratory protocol for efficient and reliable STR genotyping by sequencing with low quantity and quality DNA samples. The protocol comprises (i) an automated bioinformatic pipeline to design large sets of short tetranucleotide markers optimized for multiplex amplicon sequencing of degraded and low-template DNA; (ii) guidelines for efficient in vitro optimization of multiplex amplification using directly low quantity/quality template DNA; and (iii) a library preparation procedure that improves detection of low-level allele signal while enabling quality assessment of STR amplicon sequencing under limiting DNA conditions. We demonstrate the approach by developing and validating STR panels for non-invasive genotyping of three large carnivore species: a 44-plex for the grey wolf (Canis lupus), a 41-plex for the Eurasian lynx (Lynx lynx), and a 30-plex for the brown bear (Ursus arctos). Multiplex performance was high, with [≥]91% of samples successfully genotyped at [≥]50% of loci (allele size range 28-110 bp across panels) and correctly assigned to known individuals, negligible levels of noise in the controls, and high discriminatory power (PIDsibs [≤]2.4 x 1e-12), also owing to sequence variation among same-length alleles at 15-50% of loci. The approach is broadly applicable to animal and plant species, a wide range of sample types, and large-scale analysis such as genetic monitoring. Our study reinforces the value of STR amplicon sequencing for ecological and conservation applications while highlighting the importance of marker design and laboratory workflows tailored to HTS-based genotyping for accurate and efficient implementation.
De Keyzer, L.; Deserranno, K.; Skevin, S.; Van Hoofstat, D.; Deforce, D.; Van Nieuwerburgh, F.
Show abstract
Recombinase polymerase amplification (RPA) enables rapid nucleic acid testing in low-resource environments, but poorly characterized byproducts can compromise assay specificity and cause false-positive results. Here, we amplified the thirteen original CODIS core loci and Amelogenin to characterize recurrent RPA artefacts and establish conditions that reduce their formation. First, RPA products were analyzed for two reference samples by Oxford Nanopore Technologies sequencing. This revealed two distinct classes of multimeric products: primer multimers and amplicon multimers, consisting of repeated primer or amplicon sequences, respectively. Individual artefacts contained up to 281 primer copies or 22 amplicon copies, demonstrating the extensive range of these products. Next, we performed an optimization study to evaluate the effects of reaction temperature and reagent concentrations at two representative loci, D3S1358 and D5S818. Among the conditions tested, temperature had the most pronounced effect. Reducing the temperature from 42{degrees}C to 34{degrees}C increased the relative target amplicon fraction from 15% to 83% for D3S1358 and from 84% to 98% for D5S818, while maintaining or increasing absolute target concentration. Lower primer concentrations and higher T4 UvsX concentrations also reduced multimer formation, although lower primer concentrations reduced target yield and caused allelic dropout. Finally, amplification at 34{degrees}C was evaluated across all fourteen loci by sequencing. Relative to 42{degrees}C, the target read fraction increased by more than 5 percentage points for 7/14 loci in one reference sample and 9/14 loci in the other, with the largest improvements at multimer-prone loci. These findings identify multimers as an important class of RPA artefacts and establish reaction temperature and T4 UvsX concentration as promising conditions to improve RPA specificity.
Dewi, Y. K.; Chudori, Y. N.
Show abstract
Reliable DNA isolation is a critical prerequisite for PCR-based food authentication, particularly for meat products where complex matrices may compromise DNA quality and amplification efficiency. This study aimed to analytically validate an automated DNA extraction method from meat matrices using Qiagen QIAcube Connect in combination with the DNeasy(R) Mericon Food Kit. Validation parameters included DNA concentration, total yield, purity, integrity, and assessment of PCR inhibitors using real-time PCR targeting the porcine cytochrome b gene. The method produced a mean DNA concentration of 219.5 ng/{micro}L with an average yield of 21,519.7 ng, exceeding predefined acceptance criteria. Agarose gel electrophoresis confirmed DNA fragment sizes larger than the target amplicon, indicating suitability for PCR analysis. Real-time PCR evaluation demonstrated excellent linearity (R2 = 0.99-1.00), amplification efficiencies between 90.34% and 99.84%, and mean {Delta}Ct values of 0.10, confirming the absence of PCR inhibition. These results indicate that the validated automated method is robust, reproducible, and suitable for routine PCR-based meat species authentication in food control laboratories.
Meerson, A.
Show abstract
To explore adapting qPCR systems for end-point nucleic acid quantification using dyes such as SYTO-9, we quantified serial dilutions of DNA and RNA standards in the range of 0.75 - 200 ng/{micro}l on 384-well qPCR devices. SYTO-9 fluorescence was successfully measured using standard SYBR Green settings. Blank-subtracted relative SYTO-9 signal showed a logarithmic dependence on DNA/RNA concentration (R2 > 0.95). Measurements were highly stable with different incubation times, temperatures of up to 95{degrees}C, and photobleaching. The described approach is a valuable QC option for high-throughput DNA/RNA isolations and could be adapted to additional fluorometric assays beyond nucleic acids.
Yuan, T.; Bai, Y.; Song, L.; Zhang, A.; Liu, Y.; Cao, Y.
Show abstract
Cell-free DNA (cfDNA) methylation profiling holds great promise for non-invasive cancer detection, yet accurate methylome analysis is compromised by DNA damage inherent to cfDNA. Standard library preparation workflows involve an end-repair step during which DNA polymerases can initiate synthesis from single-strand breaks (nicks), replacing endogenous methylated nucleotides with unmethylated nucleotides in the 3' direction and systematically erasing methylation information. This artifact is distinct from the terminal jagged-end effect and disproportionately affects cfDNA and FFPE DNA, which harbor abundant nicks. Here, we developed cf-Cabernet, which builds upon the Cabernet framework (Cao et al., 2023) -- an enzymatic methylation sequencing method featuring carrier DNA-assisted sample recovery and amplification-friendly post-conversion processing -- with the addition of a Taq DNA ligase-mediated nick repair step prior to end repair. Using matched cfDNA samples, we compared cf-Cabernet against standard EM-seq and WGBS. cf-Cabernet and EM-seq both substantially outperformed WGBS in alignment rate. Critically, while standard EM-seq exhibited globally reduced methylation levels compared to WGBS, cf-Cabernet yielded methylation levels concordant with WGBS. M-bias analysis revealed that EM-seq libraries showed persistently depressed methylation across the entire read length, whereas cf-Cabernet methylation recovered to the WGBS baseline beyond the terminal [~]40 bp jagged-end region. Nick-induced methylation erasure during end repair is a significant but previously underappreciated source of error in cfDNA methylation sequencing. cf-Cabernet effectively mitigates this artifact through pre-emptive nick ligation, enabling accurate methylome profiling from damaged DNA templates. This method is broadly applicable to cfDNA, FFPE DNA, and other clinical specimens where DNA integrity is compromised, providing a robust foundation for methylation-based liquid biopsy applications.
Pastorino, B.; Touret, F.; Creton, M.; Viala, R.; Morand, J. C.; Reyre, F.; Jousserand, M.; Billecard, F.; Charrel, R. N. C.
Show abstract
The COVID-19 pandemic has imposed a reevaluation of safety protocols across various sectors, including the arts. This study addresses a critical gap in understanding SARS-CoV-2 persistence on materials commonly associated with musical instruments and scores, such as alloys, varnishes, reeds, and paper. While previous research has explored viral survival on various surfaces, limited data exists for materials specific to musical contexts. In this work, we investigate the efficacy of quarantine as a non-destructive method for inactivating SARS-CoV-2 on 16 materials, including brass, silver plating, ABS plastic, ebonite, and various varnishes and paper types. Results revealed significant variability in viral persistence across materials. Non-porous surfaces like metals and ABS plastic cleared infectivity within 3 days, while porous materials such as reeds and music scores required up to 7 days. Gold-plated brass and certain varnishes showed intermediate persistence, with infectivity clearing after 4 days. These findings are in agreement with prior studies indicating that SARS-CoV-2 survival is highly dependent on surface composition, with porous and organic-coated materials retaining viable virus longer due to reduced environmental stress. Our results highlight the feasibility of stratified quarantine protocols based on material type, offering practical guidelines for musicians and institutions and provides critical insights for mitigating SARS-CoV-2 transmission risks in musical settings.
Gan, H.; Wang, X.; Tang, F.; Ibrahim, H.; Chen, X.; Xie, P.; Zhang, S.; Lin, G.; Zeng, J.; Chu, H.; Zhang, S.
Show abstract
BackgroundNon-invasive embryo quality assessment is a critical unmet need in assisted reproductive technology (ART). Preimplantation genetic testing for aneuploidy (PGT-A) is effective but requires invasive biopsy that may compromise embryo viability. Metabolomics of spent embryo culture media (SECM) offers a non-invasive alternative, yet analytical challenges--limited sample volume, high salt content, and abundant proteins--have hindered standardization and clinical translation. ResultsWe systematically optimized sample preparation for untargeted LC-MS metabolomics of SECM using human serum as a reference. Optimal conditions were highly matrix-dependent: SECM required 7x volume of 50% acetonitrile for extraction and 40% acetonitrile for reconstitution, whereas serum required 10x volume of 100% methanol and 100% water--reflecting that SECM contains more non-polar species than serum. Applying the optimized workflow to 120 clinical SECM samples (72 euploid, 48 aneuploid), we identified 102 differential metabolites between euploid and aneuploid embryos, with prominent enrichment of lipid pathways (fatty acid metabolism, {beta}-oxidation, sphingolipid metabolism) and involvement of amino acid (methionine, tryptophan) and TCA cycle metabolism. A weighted ensemble machine learning model discriminated aneuploid from euploid embryos with an AUC of 0.977, 100.0% specificity, and 89.6% sensitivity. Among 72 euploid embryos stratified by morphological grading (good, fair, poor), metabolic alterations progressed from mitochondrial energy deficiency (good vs. fair) to broader lipid dysregulation (fair vs. poor), with the ensemble model achieving AUCs of 0.944, 0.889, and 0.943, respectively. ConclusionsThis study establishes a rigorously optimized and validated SECM metabolomics workflow that overcomes key analytical barriers in this challenging matrix. Our findings demonstrate that metabolic signatures--particularly in lipid and energy metabolism--are strongly associated with both embryo ploidy and morphological quality, providing biological insights into the metabolic underpinnings of embryo developmental competence. The high predictive performance of the ensemble model supports the feasibility of non-invasive embryo assessment as a complementary tool to existing methods, with potential to reduce reliance on invasive biopsy in ART. External validation in prospective multi-center cohorts is warranted to further assess clinical utility and generalizability. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=122 SRC="FIGDIR/small/741666v1_ufig1.gif" ALT="Figure 1"> View larger version (59K): org.highwire.dtl.DTLVardef@1be208aorg.highwire.dtl.DTLVardef@14a61a3org.highwire.dtl.DTLVardef@504e70org.highwire.dtl.DTLVardef@4dc12d_HPS_FORMAT_FIGEXP M_FIG C_FIG
Engels, I.; Dedrie, T.; Saugen, S. M.; Van de Vyver, S.; Vandenbroucke, T.; Di Modica, K.; Decher, J.; Toso, A.; Deforce, D.; Daled, S.; Burnett, A.; Abrams, G.; Dhaenens, M.
Show abstract
Species identification in palaeoproteomics relies on genome-derived protein sequences which are often poor-quality, and lacks tools to cope with multi-species samples. Here, we address both challenges through the analysis of physical and genetic mixtures. Species that are absent from our database are considered a genetic mixture, i.e. a patchwork of peptides from closely related species. Inversely, various overlapping peptide stretches allow us to resolve complex physical mixtures. This is benchmarked by analysing physical mixtures of modern bone fragments, including genetic mixtures. We illustrate the impact of our approach via a rapid and high-throughput analysis of >2500 bone fragments, revealing the Eemian-era faunal environment around Scladina Cave, including the first Palaeoloxodon antiquus identified at this site. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=182 SRC="FIGDIR/small/732552v1_ufig1.gif" ALT="Figure 1"> View larger version (67K): org.highwire.dtl.DTLVardef@75c1d3org.highwire.dtl.DTLVardef@1084481org.highwire.dtl.DTLVardef@1c9a4f9org.highwire.dtl.DTLVardef@16dd859_HPS_FORMAT_FIGEXP M_FIG C_FIG
Elander, B. E.; Jiang, M.; Guthrie, C.; Ibrahim, Z.; Momeni, B.; Wang, D.
Show abstract
In response to the growing need to recycle E-waste, specifically that of lithium-ion batteries (LIBs), the development of sustainable recycling methodologies will be vital. As one of the most sustainable and low-cost options, biohydrometallurgy (BioHM) uses biological organisms to facilitate the recovery of critical metals from spent LIBs. Despite the advantages of BioHM, its slow kinetics, due to the reliance on the metabolic activity of microorganisms, limit its large-scale application for the closed-loop recycling of LIBs. In this work, we investigate the correlation between the incubation time of Acidithiobacillus ferrooxidans (Atf) and its leaching efficiency of four common elements in LIBs: Li, Ni, Mn, and Co. We assess how specific incubation times along the biological growth curve affect the leaching efficiency of Li[Ni0.6Mn0.2Co0.2]O2 (NMC622), a model electrode material. Our results show that pH alone is not an accurate descriptor of the leaching efficiency of Atf cultures used at different growth stages. The addition of NMC622 during the bacterial lag phase reaches similar or even improved extraction compared to cultures used after reaching the exponential or stationary phases. A simple model of leaching dynamics shows predictions consistent with our experimental observations under the condition that the inhibition of bacterial growth by NMC is not severe. Our findings indicate that early-stage cultures can alleviate the kinetic bottleneck and improve the throughput of recovering critical materials from spent batteries.
Apweiler, M.; Broche, J.; Loitz, M.; Hackenbruch, L.; Ossowski, S.; Schroeder, C.; Schmit, K. J.
Show abstract
Cell-free DNA (cfDNA), released from apoptotic and necrotic cells into body fluids, represents a non-invasive source of genetic information for disease prediction, diagnosis, and monitoring. However, its low physiological abundance makes cfDNA highly susceptible to pre-analytical influences. In particular, genomic DNA (gDNA) released from lysed white blood cells (WBCs) can contaminate plasma and compromise downstream cfDNA analyses. This study evaluated the impact of different blood collection tubes and isolation methods on cfDNA stability and yield. Blood samples from 13 healthy donors were collected using cfDNA-stabilizing tubes (Cell-Free DNA BCT, Streck; S-Monovette cfDNA Exact, Sarstedt) and stored at room temperature for 1, 5, or 10 days before plasma isolation. CfDNA was extracted using either a magnetic bead-based method or a silica column-based approach. DNA quantity and quality were assessed by fluorometric quantification, automated fragment analysis, and gene-specific quantitative PCR. Streck-based workflows maintained stable cfDNA yields and characteristic mononucleosomal fragmentation profiles across all storage times. In contrast, Sarstedt tubes showed reduced cfDNA concentrations after 5 days and a pronounced increase at 10 Days, accompanied by high-molecular weight DNA patterns consistent with WBC lysis. These trends were largely independent of the extraction method. Overall, the results demonstrate that blood collection tube chemistry critically influences cfDNA integrity during delayed processing. Streck tubes, particularly when combined with QIAamp, provided the most robust and reproducible workflow for routine molecular diagnostics, whereas Sarstedt tubes produced physiologically implausible results after extended storage.
Kobara, S.; Huang, M.; Ilhamsyah, R.; Struk, D.; Dimandja, J.; Hesketh, P. J.; Arrubla, D. C.; Fensore, C.; Polito, C.; Kamaleswaran, R.; Esper, A.
Show abstract
Background/ObjectivesPolydimethylsiloxane (PDMS) is a non-invasive and versatile material often used for non-invasive collection of skin-emitted volatile organic compounds (VOCs), with potential applicability in acute and pre-critical care settings. However, most existing PDMS-based methodologies rely on extensive sample preparation and environmental control, limiting their feasibility in time-sensitive clinical contexts. MethodsWe conducted a proof-of-concept pilot study in four healthy volunteers to evaluate whether a simplified skin-contact PDMS sampling procedure can capture detectable VOCs and preserve individual-level variation. PDMS strips were applied directly to the skin with minimal preparation, and collected VOCs were analyzed using gas chromatography-mass spectrometry. Donor-associated variability was assessed using Bray-Curtis dissimilarity, and variability in VOC detection was evaluated across body sites. ResultsSkin-contact PDMS sampling detected 160 VOCs across four participants. The mean within-donor Bray-Curtis dissimilarity was 0.308, compared with a mean between-donor dissimilarity of 0.347. Preliminary permutation testing showed distinguishable donor profiles (p-value = 0.004). VOC detection variability differed across body sites, with lower coefficients of variation at the forehead, neck, and wrist than at the ankle. ConclusionsUnder simplified sampling conditions, skin-contact PDMS captured individual-associated VOC profiles with lower within-donor variability than between-donor variability. These findings support the feasibility of PDMS-based skin VOC sampling in minimally controlled settings. Further validation in larger and clinically relevant cohorts is warranted to assess the utility of PDMS-sampled skin VOCs as potential biomarkers for early disease detection.
Płonka, W.; Kostka, D.; Lalik, A.; Kurpas, M.; Dinh, K. N.; Sitkiewicz, M.; Kimmel, M.; Rzyman, W.; Jaksik, R.
Show abstract
Formalin-fixed, paraffin-embedded (FFPE) tissues remain an essential resource for molecular studies, yet formalin-induced cytosine deamination introduces characteristic C>T/G>A artifacts that compromise the accuracy of next-generation sequencing (NGS) analyses. Numerous computational methods and enzymatic DNA repair strategies have been proposed to reduce these artifacts, but no systematic comparison across tools and experimental conditions exists. Here, we evaluate the performance of seven computational approaches (SOBDetector, Ideafix, MicroSEC, FFPolish, DeepOmics FFPE/FFPE-PLUS, FFPErase) together with the NEBNext(R) FFPE DNA Repair Mix v2, a multi-enzyme repair system applied during DNA preparation. Using three independent datasets, one based on whole genome sequencing (CGCI-BL) and two on whole exome sequencing (TCGA-PC and SUT-LUAD, the latter containing enzymatically repaired samples), and matched fresh-frozen samples as the gold standard, we assess precision, sensitivity, and artifact reduction efficiency across all methods. We further examine the potential synergy between enzymatic repair and post-sequencing computational filtering. Our results provide practical guidelines for FFPE artifact correction and demonstrate that enzymatic treatment provides the best results, while among the computational methods, FFPErase offers the most robust reduction of cytosine deamination artifacts while maximizing the retention of true somatic variants. KEY MESSAGESO_LIFormalin fixation in FFPE samples introduces artifacts that can significantly affect the accuracy of NGS analyses. C_LIO_LIAmong the evaluated approaches, enzymatic repair using NEBNext(R) FFPE DNA Repair Mix v2 achieves the most effective reduction of sequencing artifacts. C_LIO_LIComputational methods vary in performance, with FFPErase showing the most robust balance between artifact removal and retention of true somatic variants. C_LIO_LICombining enzymatic repair with computational filtering did not lead to consistent improvements in performance across datasets. C_LI
Zhang, Z.; Xu, Y.
Show abstract
This study aims to quantify the genetic similarity of different species (from fish to humans) to the human reference genome (pp6, Homo sapiens.GRCh38) based on the allele presence/absence patterns of 33 language/cognition related gene SNV loci, identify key breakpoints during evolution, and evaluate the enrichment of language and cognition genes at these breakpoints. We designed a similarity calculation method relying on binary features (four columns for A/T/C/G), adopted five difference/distance measures (Sorensen, Rogers, Nei, Reynolds, and Hellinger), and converted them into similarity values (1/(1+distance)). For each method, samples were independently ranked, the first derivative of similarity was computed, and the top 12 peaks were selected as candidate breakpoints. Results show that the similarity curves from the five methods are highly consistent (correlation coefficients >0.9), with major peaks concentrated at positions 355, 363, 381, 382, 390, 400, etc., where the corresponding samples are predominantly ancient hominins and primates. Furthermore, we defined 13 peak groups (starting positions 355-401). For each peak within a group, pairwise SNV differences between the peak apex sample and its immediate left neighbor were compared, and the intersection F_INTERSECTION (shared differential loci) was obtained. For each F_INTERSECTION, we calculated the proportions of language genes and cognition genes. In addition, we computed the differential sets between adjacent groups' F_INTERSECTION to trace the gradual emergence of new loci. In F_INTERSECTION, language genes accounted for an average of 59.5%, and cognition genes for an average of 62.9%. The proportion of language genes reached a peak at position 383 (61.2%), while cognition genes peaked at position 386 (64.9%). High frequency peak samples include c25, c27, and ja2, suggesting that language cognition genes may have undergone independent intensification during Eurasian evolution. Differential analysis between adjacent F_INTERSECTION revealed a stepwise acquisition of new loci from position 355 to 401, with three bursts of newly added loci along the entire evolutionary axis. This study provides a quantitative framework based on similarity curves, offers a novel molecular perspective for understanding the evolution of language and cognitive abilities, and highlights the potential importance of East Asian archaic hominins in the evolution of language cognition genes.
Watson, E.; Qian, G.; Ravishankar, S.; Hobbs, M.; Copty, J.; Yu, C.; Kummerfeld, S.; Liang, C.; Lacaze, P.; Davis, R. L.; Sue, C. M.
Show abstract
Mitochondrial diseases (MDs) are clinically heterogeneous rendering ascertainment challenging. Estimates of pathogenic mitochondrial DNA (mtDNA) variants in the population range from 1 in 200 to 1 in 4,000 individuals. Inclusion of mtDNA sequencing in genomic databases facilitates comprehensive estimation of mtDNA variation. However, interpretation of low heteroplasmy variation is complex, due in part to misalignment of nuclear mitochondrial DNA transcripts (NUMTs), whilst conservative heteroplasmy thresholds likely omit relevant variation. Cumulative burden of mtDNA variation contributes to aging and neurodegeneration, and recent characterisation of mitochondrial genome constraint allows quantitation of this burden. We analysed whole genome sequencing of blood DNA from 3,500 healthy older individuals in the Medical Genome Reference Bank using mity, considering pathogenic mtDNA variants [≥]1% heteroplasmy. We identified 34 distinct pathogenic mtDNA variants in 62 individuals, giving a combined population allele frequency of 1.77% (95% CI 1.36-2.27) or 1 in 56 individuals. We evaluated inclusion of false positive (FP) calls due to two common NUMTs, which accounted for up to 16% of variants. Increasing heteroplasmy thresholding to eliminate all NUMT-FPs also eliminated much of the total variation, including pathogenic variants. We propose a sample-specific, scaled heteroplasmy threshold to maximise variant retention and mitigate NUMT-FPs. Finally, we characterised measures of mitochondrial constraint in this healthy older cohort, observing an association between variant burden and summed constraint, whilst mean constraint was higher in pathogenic variant carriers. These findings suggest pathogenic mtDNA variation is more common in the population than is currently appreciated. Findings are comparable to larger genomic databases when heteroplasmy thresholding is adjusted, and support earlier population-based estimates. Incorporation of low heteroplasmy variation is relevant, but interpretation is nuanced, and optimising variant retention requires consideration of NUMT-FP rates.